Designed for modern deep neural networks that use ReLU,

W𝒩(0,2nl)W \sim \mathcal{N}\left(0,\frac{2}{n^l}\right)

Target: ensure activation variance across different layers

Assumptions: ReLU activation, weight normally distributed with mean of zero, weight and activations are independent.


References

  1. He, Kaiming, et al. "Delving deep into rectifiers: Surpassing human-level performance on imagenet classification." Proceedings of the IEEE international conference on computer vision. 2015.